Back

Journal of Clinical Epidemiology

Elsevier BV

All preprints, ranked by how well they match Journal of Clinical Epidemiology's content profile, based on 31 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Most clinical trials involving American children that violated FDAAA legal reporting requirements had not published outcomes in the scientific literature

Bruckner, T.; Yerunkar, S.; Basegmez, O.; Borana, R.; Chavarria, B.; Velarde, M.; DeVito, N. J.

2023-01-18 pediatrics 10.1101/2023.01.17.23284683 medRxiv
Top 0.1%
58.4%
Show abstract

Non-publication, incomplete publication and excessively slow publication of clinical trial outcomes contribute to research waste and can harm patients. In this cohort study, we used public ClinicalTrials.gov registry records to identify 81 paediatric clinical trials in the United States that appear to have violated the FDA Amendments Act 2007 (FDAAA) reporting requirements. We then searched the literature for the outcomes of these 81 trials and contacted trial sponsors about the status of results. We found that only 22/81 trials (27.2%) had made their results public in a full-length peer-reviewed publication, and that only 8/81 (9.9%) trials had done so within one year of their primary completion date. Our findings highlight the need for US Food and Drug Administration and the National Institutes of Health to systematically monitor FDAAA compliance and enforce reporting requirements.

2
Estimating the prevalence of discrepancies between study registrations and publications: A systematic review and meta-analyses

TARG Meta-Research Group & Collaborators, ; Thibault, R. T.; Clark, R.; Pedder, H.; van den Akker, O.; Westwood, S.; Munafo, M.

2021-07-28 health systems and quality improvement 10.1101/2021.07.07.21259868 medRxiv
Top 0.1%
54.4%
Show abstract

ObjectivesProspectively registering study plans in a permanent time-stamped and publicly accessible document is becoming more common across disciplines and aims to reduce risk of bias and make risk of bias transparent. Selective reporting persists, however, when researchers deviate from their registered plans without disclosure. This systematic review aimed to estimate the prevalence of undisclosed discrepancies between prospectively registered study plans and their associated publication. We further aimed to identify the research disciplines where these discrepancies have been observed, whether interventions to reduce discrepancies have been conducted, and gaps in the literature. DesignSystematic review and meta-analyses. Data sourcesScopus and Web of Knowledge, published up to 15 December 2019. Eligibility criteriaArticles that included quantitative data about discrepancies between registrations or study protocols and their associated publications. Data extraction and synthesisEach included article was independently coded by two reviewers using a coding form designed for this review (osf.io/728ys). We used random-effects meta-analyses to synthesize the results. ResultsWe reviewed k = 89 articles, which included k = 70 that reported on primary outcome discrepancies from n = 6314 studies and, k = 22 that reported on secondary outcome discrepancies from n = 1436 studies. Meta-analyses indicated that between 29% to 37% (95% confidence interval) of studies contained at least one primary outcome discrepancy and between 50% to 75% (95% confidence interval) contained at least one secondary outcome discrepancy. Almost all articles assessed clinical literature, and there was considerable heterogeneity. We identified only one article that attempted to correct discrepancies. ConclusionsMany articles did not include information on whether discrepancies were disclosed, which version of a registration they compared publications to, and whether the registration was prospective. Thus, our estimates represent discrepancies broadly, rather than our target of undisclosed discrepancies between prospectively registered study plans and their associated publications. Discrepancies are common and reduce the trustworthiness of medical research. Interventions to reduce discrepancies could prove valuable. Registrationosf.io/ktmdg. Protocol amendments are listed in Supplementary Material A.

3
Retracted randomized trials attributed to super-retractors and top-cited scientists with multiple retractions: secondary analysis of the VITALITY retrospective cohort

Lyu, C.; Matbouriahi, M.; Naudet, F.; Ioannidis, J. P. A.; Cristea, I. A.

2025-11-25 epidemiology 10.1101/2025.11.23.25340834 medRxiv
Top 0.1%
52.4%
Show abstract

ImportanceMultiple retractions from the same author often uncover issues affecting their entire work, such as having systematically altered or fabricated data. ObjectivesEvaluate the contribution of authors with most retractions ("super-retractors") and top-cited scientists with multiple retractions to the retracted clinical trial literature. DesignRetrospective cohort study, linking an openly available cohort (VITALITY) of 1330 retracted randomized clinical trials (RCTs) to three lists of scientists: super-retractors, totaling most retractions in the Retraction Watch Leaderboard, and top-cited scientists, over the entire career or in the most recent single year, who accumulated 10 or more retractions not due to editor/publisher errors. The VITALITY cohort was updated up to November 2024. The three author lists were updated in August 2025. Participants30 super-retractors, 163 career-long and 174 single-year scientists totaling 10 or more retractions. Main outcomesAuthorship and characteristics of retracted RCTs (publication and retraction year, time between publication and retraction, number of citations). Results6/30 super-retractors, representing Anesthesiology and Endocrinology & Metabolism, co-authored 290/1330 retracted RCTs (22%). 18/163 career-long top-cited scientists with at least 10 retractions, representing 10 fields, co-authored 327/1330 trials (25%), 275 (84%) of which were also co-authored by a super-retractor. 7/174 single-year top-cited scientists with at least 10 retractions co-authored 50 retracted trials; all of them were also among the career-long top-cited scientists with at least 10 retractions. Articles with super-retractors authors vs not were published earlier (median (IQR)= 2000 (1997-2005) vs 2020 (2014-2022)); retracted earlier (median (IQR)= 2013 (2012-2019) vs 2023 (2018.5-2023)); had a longer lag between publication and retraction, (median (IQR)= 5111 (3560-6820) vs 482 (330-1119) days); and accrued more citations (median (IQR)= 21 (12-42) vs 5 (1-19)). In multivariable regression models, only time to retraction ({beta} = 0.02, P < 0.001) was significantly and positively associated with total citations. Results were similar when comparing retracted articles from top-cited scientists with at least 10 retractions versus other articles. Conclusions and relevanceIn this cohort study of 1330 retracted RCTs, a small number of influential authors, often co-authors and concentrated across few fields of medicine and countries, account for a significant proportion of retracted clinical trials. Key pointsO_ST_ABSQuestionC_ST_ABSWhat is the contribution of the authors with most retractions ("super-retractors") and of those top-cited with multiple retractions to the retracted randomized clinical trials literature? FindingsIn this cohort study, six super-retractors, from Anesthesiology and Endocrinology & Metabolism, co-authored one fifth of all retracted trials, while 18 top-cited scientists with over 10 retractions co-authored a quarter of them. Articles co-authored by super-retractors or by top-cited scientists with multiple retractions were published and retracted earlier, took longer to retract and accumulated more citations. MeaningRetracted clinical trials are disproportionately associated with a small number of influential authors, often co-authors and concentrated across few subfields of medicine and countries.

4
Assessment of Bias in Clinical Trials with LLMs Using ROBUST-RCT: A Feasibility Study.

Vidor, P. R.; Casiraghi, Y.; de Souza, A. M.; Schmidt, M. I.

2025-08-13 epidemiology 10.1101/2025.08.12.25333520 medRxiv
Top 0.1%
45.7%
Show abstract

BACKGROUNDBias assessment is a crucial step in evaluating evidence from randomized controlled trials. The widely adopted Cochrane RoB 2, designed to identify these issues, is complex, resource-intensive, and unreliable. Advances in artificial intelligence (AI), particularly in the field of large language models (LLMs), now allow the automation of complex tasks. While prior investigations have focused on whether LLMs could perform assessments with RoB 2, integrating technologies does not resolve the intrinsic methodological issues of the instrument. This is the first feasibility study to evaluate the reliability of ROBUST-RCT, a novel bias assessment tool, as applied by humans and LLMs. METHODSA sample of RCTs of drug interventions was screened for eligibility. Reviewers working independently used ROBUST-RCT to assess different aspects of the studies and then reached a consensus through discussion. A chain-of-thought prompt instructed four LLMs on how to apply ROBUST-RCT. The primary analysis used Gwets AC2 coefficient and benchmarking to assess inter-rater reliability of the "judgment set", defined as the series of final assessments for the six core items in the ROBUST-RCT tool. RESULTS54 assessments of each LLM were compared to human consensus in the primary analysis. Gwets AC2 inter-rater reliability ranged from 0.46 to 0.69. With 95% confidence, three of the four tested LLMs achieved moderate or higher reliability based on probabilistic benchmarking. A secondary analysis also found a Fleiss Kappa of 0.49 (95% CI: 0.30 - 0.60) between human reviewers before consensus, numerically higher than the values reported in prior literature about RoB 2. CONCLUSIONLarge Language Models (LLMs) can effectively perform risk-of-bias assessments using the ROBUST-RCT tool, enabling their integration into future systematic review workflows aiming for enhanced objectivity and efficiency.

5
Trustworthiness and Transparency Features Were Less Frequent in Randomized Trials Presenting Large Effects in Abstracts

Henssler, J.; Reis-Pardal, J.; Koppel, L.; Ioannidis, J.

2025-10-22 epidemiology 10.1101/2025.10.20.25338369 medRxiv
Top 0.1%
45.6%
Show abstract

OBJECTIVESLarge effect sizes (ESs), especially when prominently presented in trial abstracts, draw large attention, but it is important to understand whether they are trustworthy. We aimed to assess indicators of transparency and trustworthiness in randomized controlled trials (RCTs) reporting some large ES in their abstract, in comparison with RCTs presenting only non-large ESs in their abstract. STUDY DESIGN AND SETTINGWe included RCTs indexed in MEDLINE between January 1, 2024 and March 18, 2025, presenting at least one standardized mean differences of absolute value 0.8 or higher (large ES) versus those presenting only smaller absolute standardized mean differences in their abstract. Trial characteristics and methodological features were extracted systematically in large ES and non-large ES trials. Primary outcome was pre-specified protocol registration, secondary outcomes were having no protocol and public availability or repository placement of raw data. RESULTSWe evaluated 152 trials with large ESs in their abstract and 175 trials with only non-large ESs in their abstract. Large ES trials had suggestively lower rates of pre-registered protocols (45% versus 61%, p=0.0054) and significantly higher rates of no protocol registration (26% versus 13%, p=0.0028) than non-large ES trials. There was no difference in raw data public availability or repository placement (6% versus 7%). Large ES trials were also less likely to be multicenter (p=0.0042), to have high-income country of corresponding author (p=0.0001), to be conducted in high-income country site(s) (p=0.0003), to have a published statistical analysis plan (p=0.0216), and to result from between-group comparisons (p<0.0001). Large effects were significantly more likely to involve non-drug/non-psychological interventions (p=0.0001). CONCLUSIONSRCTs presenting large ESs in their abstracts are more likely to lack transparency and trustworthiness features and may operate with higher risk of lack of credibility. HighlightsO_LIThis meta-research study assessed the association of effect sizes with features of transparency and trustworthiness of RCTs C_LIO_LILarge effect trials have fewer registered and available protocols C_LIO_LIStudies with large effects sizes are more likely to have a smaller sample size, single center recruitment and provenance from countries without strong clinical research tradition C_LIO_LIData sharing is poor across all trials regardless of effect size C_LI What is new?O_ST_ABSKey findingsC_ST_ABSO_LIThis meta-research study indicated that RCTs presenting large effect sizes in their abstracts are more likely to lack features of transparency and trustworthiness. C_LI What this adds to what is known?O_LILarge effect sizes in abstracts of trials draw attention and may seem as compelling evidence for a studys claim. Yet there have been divergent findings concerning the association between effect size and credibility of a trial. Our study showed that large effect sizes are consistently associated with a range of methodological shortcomings. C_LI What is the implication and what should change now?O_LIRCTs presenting large effect sizes may operate with higher risk of lack of credibility. C_LIO_LIRather than indicating a convincing finding, extreme results should be met with increased attention and careful scrutiny. C_LI

6
State of play in individual participant data meta-analyses of randomised trials: Systematic review and consensus-based recommendations

Seidler, A. L.; Aagerup, J.; Nicholson, L.; Hunter, K.; Bajpai, R.; Hamilton, D.; Love, T.; Marlin, N.; Nguyen, D.; Riley, R.; Rydzewska, L.; Simmonds, M.; Stewart, L.; Tam, W.; Tierney, J.; Wang, R.; Amstutz, A.; Briel, M.; Burdett, S.; Ensor, J.; Hattle, M.; Libesman, S.; Liu, Y.; Schandelmaier, S.; Siegel, L.; Snell, K.; Sotiropoulos, J.; Vale, C.; White, I.; Williams, J.; Godolphin, P.

2026-02-04 epidemiology 10.64898/2026.02.03.26345481 medRxiv
Top 0.1%
40.8%
Show abstract

BackgroundIndividual participant data (IPD) meta-analyses obtain, harmonise and synthesise the raw individual-level data from multiple studies, and are increasingly important in an era of data sharing and personalised medicine to inform clinical practice and policy. Objectives(1) Describe the landscape of IPD meta-analysis of randomised trials over time; (2) establish current practice in design, conduct, analysis and reporting for pairwise IPD meta-analysis; and (3) derive recommendations to improve the conduct of and methods for future IPD meta-analyses. DesignPart 1: systematic review of all published IPD meta-analyses of randomised trials; Part 2: in-depth review of current methodological practice for pairwise IPD meta-analysis; and Part 3: adapted nominal group technique to derive consensus recommendations for IPD meta-analysis authors, educators and methodologists. Data sourcesMEDLINE, Embase, and the Cochrane Database of Systematic Reviews (via the Ovid interface). Eligibility criteriaPart 1: all IPD meta-analyses of randomised trials published before February 2024, evaluating intervention effects and based on a systematic search. Part 2: all pairwise IPD meta-analyses from part 1 published between February 2022 and February 2024. Part 3: Selected panel of experienced IPD meta-analysis authors and/or methodologists. ResultsPart 1: We identified 605 eligible IPD meta-analyses published between 1991 and 2024. The number of IPD meta-analyses published per year increased over time until 2019 but has since plateaued to about 60 per year. The most common clinical areas studied were cardiovascular disease (n=113, 19%) and cancer (n=110, 18%). The proportion of IPD meta-analyses published with Cochrane decreased over time from 16% (n=31/196) before 2015 to 3% (n=5/196) between 2021-2024. Part 2: 100 recent pairwise IPD meta-analyses were included in the in-depth review. Most cited PRISMA-IPD (68, 68%) and conducted risk of bias assessments (n=82, 82%), with just under half carrying out subgroup analyses not at risk of aggregation bias (n=36/85, 41%). However, only 33% (n=33) and 29% (n=29) respectively provided a protocol or statistical analysis plan, and only 7% (n=6/82) reported using IPD to inform risk of bias assessments. Part 3: 24 experts participated in a consensus workshop. Key recommendations for improved IPD meta-analyses focused on transparency (prospective registration; published protocols and statistical analysis plans) and maximising value (searching trial registries; obtaining IPD for unpublished evidence; using IPD to address missing data and risk of bias). Methodologists and educators should strengthen dissemination of methods and support capacity building across clinical fields and geographical areas. ConclusionsThe application and methodological quality of IPD meta-analyses of randomised trials has increased in the last decade, but shortcomings remain. Implementing our consensus-based recommendations will ensure future IPD meta-analyses generate better evidence for clinical decision making. Study registrationOpen Science Framework (1) Summary boxesO_ST_ABSWhat is already known on this topicC_ST_ABSO_LIIPD meta-analyses of randomised trials are regularly used to inform clinical policy and practice. C_LIO_LIThey can provide better quality data and enable more thorough and robust analyses than standard aggregate data meta-analyses, but are resource-intensive and can be challenging to conduct, leading to variable methodological quality C_LIO_LIPrevious studies that evaluated the conduct of IPD meta-analyses pre-date several major developments, such as the introduction of the PRISMA-IPD reporting guideline. C_LI What this study addsO_LIThis is the most comprehensive assessment of IPD meta-analyses of randomised trials to date (605 studies), showing an increase in publications over time followed by a recent plateau. C_LIO_LIThe conduct of IPD meta-analysis has improved in recent years including increased use of prospective registration, assessment of risk of bias, appropriate analyses of patient subgroup effects and citing the PRISMA-IPD statement. C_LIO_LIMany shortcomings remain including (i) insufficient pre-specification of methods such as outcomes and analyses, (ii) sub-standard transparency (including publication of protocols, statistical analysis plans and reporting of analyses), and (iii) failure to gain maximum value of IPD (i.e. include unpublished trials, use the IPD to inform risk of bias and trustworthiness assessments, and address missing data appropriately); expert consensus recommendations are provided for how to address these gaps. C_LI

7
Comparision of Bayesian methods for incorporating adult clinical trial data to improve certainty of treatment effect estimates in children.

Walker, R.; Dias, S.; Phillips, B.

2023-02-02 pediatrics 10.1101/2023.02.02.23285367 medRxiv
Top 0.1%
40.0%
Show abstract

There are challenges associated with recruiting children to take part in randomised clinical trials and as a result, compared to adults, in many disease areas we are less certain about which treatments are most safe and effective. This can lead to weaker recommendations about which treatments to prescribe in practice. However, it may be possible to borrow strength from adult evidence to improve our understanding of which treatments work best in children, and many different statistical methods are available to conduct these analyses. In this paper we discuss Bayesian methods for extrapolating adult clinical trial evidence to children. Using an exemplar dataset, we compare the effect of modelling assumptions on the estimated treatment effect and associated heterogeneity. We finally discuss the appropriateness of different modelling assumptions in the context of estimating treatment effect in children.

8
A Methodological Evaluation of Meta-Analyses in tDCS - Motor Learning Research

Alsalti, T.; Hussey, I.; Elson, M.; Krause, R.; Pohl, S.

2024-07-27 rehabilitation medicine and physical therapy 10.1101/2024.07.26.24311068 medRxiv
Top 0.1%
38.7%
Show abstract

With transcranial direct-current stimulations (tDCS) rising popularity both in motor learning research and as a commercial product, it is becoming increasingly important that the quality of evidence on its effectiveness be evaluated. Special attention should be paid to meta-analyses, as they usually have a large impact on research and clinical practice. The aim of this study was to evaluate the methodological quality of meta-analyses estimating the effect of tDCS on motor learning with respect to reproducibility as the main focus, and reporting quality and publication bias control as secondary aspects. The three meta-analyses we reviewed largely adhered to PRISMA reporting guidelines and reported the primary effect sizes and sampling variances / confidence intervals they calculated, enabling successful reproductions of pooled effect size estimates. However, akin to previous meta-research reviews with similar aims, we found the methods and results sections of the meta-analyses to be severely underreported, which compromises the ability to judge the soundness of the methodological procedure adopted as well as its reproducibility. While publication bias detection methods were applied, the approaches chosen do not allow for well informed decisions about the presence or extent of publication bias. These results reemphasise the need to transparently report methods in meta-analyses and to meticulously evaluate their quality before and after publication.

9
LLM-assisted evidence audit of late-stage cancer incidence as a screening trial endpoint

Li, S.; Zhang, W.; Xing, X.; Shen, Z.; Wang, Y.; Chen, Z.; Neto, O.; Yu, Y.; Wu, C.; Lin, L.

2026-08-31 oncology 10.64898/2026.08.29.26361733 medRxiv
Top 0.1%
38.6%
Show abstract

Background Late-stage cancer incidence is being considered as an earlier endpoint in cancer-screening trials, but its trial-level association with cancer-specific mortality may depend on evidence selection and endpoint harmonization. We evaluated the robustness of this association to source-verified additions. Methods We reconstructed the PubMed corpus underlying a 41-comparison review. Gemini 3.1 Pro Preview was used only to prioritize reports for blinded human reassessment. Reviewers determined eligibility, linked reports from the same trial, harmonized endpoints, and verified comparison-level data. We recalculated unweighted Pearson correlations overall and by cancer type after adding earliest-compatible trial comparisons. Results Among 1209 candidate records, 996 PDFs were assessed. Thirty-three reports absent from the source review were prioritized; 26 were eligible, representing 18 trials, and 8 provided compatible comparisons. Adding these comparisons increased the dataset from 41 to 49 and attenuated the overall correlation from 0.73 (95% confidence interval [CI] = 0.55 to 0.85) to 0.59 (95% CI = 0.37 to 0.75). Updated correlations were 0.49 (95% CI = -0.26 to 0.87) for breast, -0.23 (95% CI = -0.71 to 0.40) for colorectal, and 0.83 (95% CI = 0.54 to 0.95) for lung cancer. One sparse-event comparison influenced the colorectal estimate. Conclusions The overall association was sensitive to evidence composition, and cancer-specific stability varied. Late-stage incidence should be evaluated by cancer type and with prespecified sensitivity analyses for evidence selection and endpoint definitions. Model-assisted prioritization cannot replace human eligibility review, trial reconciliation, and source verification.

10
Time-to-retraction and likelihood of evidence contamination (VITALITY Extension I): a retrospective cohort analysis

Yuan, Y.; Peng, Z.; Doi, S. A. R.; Furuya-Kanamori, L.; Cao, H.; Lin, L.; Chu, H.; Loke, Y.; Mol, B. W.; Golder, S.; Vohra, S.; Xu, C.

2026-02-24 epidemiology 10.64898/2026.02.20.26346631 medRxiv
Top 0.1%
38.6%
Show abstract

BackgroundThe number of problematic randomized clinical trials (RCTs) has risen sharply in recent decades, posing serious challenges to the integrity of the healthcare evidence ecosystem. ObjectiveTo investigate whether retraction of problematic RCTs could reduce evidence contamination. DesignRetrospective cohort study SettingA secondary analysis of the VITALITY Study database. Participants1,330 retracted RCTs with 847 systematic reviews. MeasurementsThe difference in the median number (and its interquartile, IQR) of contamination before and after retraction. The association between time-to-retraction and likelihood of evidence contamination. ResultsAmong these retracted RCTs, 426 led to evidence contamination, resulting in 1,106 contamination events (251 after retraction vs. 855 before retraction). The time interval between RCT publication and first contamination ranged from 0.2 to 30.9 years, with a median of 3.3 years (95% CI: 3.0 to 3.9). The median number of contaminated systematic reviews was lower after retraction than before retraction (0, IQR: 0 to 1 vs. 1, IQR: 1 to 2, P < 0.01). Compared with trials retracted more than 7.5 years after publication, those retracted between 1.0 and 1.8 years (OR = 0.70, 95% CI: 0.60 to 0.80) and retracted within 1.0 year (OR = 0.69, 95% CI: 0.60 to 0.80) were associated with lower likelihood of evidence contamination. LimitationsOnly assessed contaminated systematic reviews with quantitative synthesis and limited to retracted RCTs. ConclusionsRetracting problematic RCTs can significantly reduce evidence contamination, and faster retraction was associated with less contamination. To safeguard the integrity of the evidence ecosystem, academic journals should act promptly in the retraction of problematic studies to minimize their downstream impact. Primary Funding SourcesThe National Natural Science Foundation of China (72204003, 72574229)

11
Quality of systematic reviews on physiotherapy interventions for musculoskeletal disorders is critically low: a meta-epidemiological study

Ferri, N.; Ravizzotti, E.; Bracci, A.; Carreras, G.; Pillastrini, P.; Di Bari, M.

2023-11-20 rehabilitation medicine and physical therapy 10.1101/2023.11.20.23298761 medRxiv
Top 0.1%
34.9%
Show abstract

QuestionHow good is the quality of systematic reviews on the effectiveness of physiotherapy for musculoskeletal conditions? Are there any factors associated with quality? DesignThis is a meta-epidemiological study on systematic reviews with meta-analysis (SR-MA) of randomised controlled trials (RCT). MethodsMEDLINE, Cochrane Database of Systematic Reviews (CDSR), CINAHL, and PEDro were searched for SR-MA of RCT on physiotherapy in musculoskeletal disorders in the last ten years. Two independent researchers screened and extracted the records and analysed the full-texts. The quality of SR-MA was quantified with AMSTAR-2 tool on a sample of 100 studies, randomly selected from the records retrieved. Disagreements were solved by consensus. ResultsThe number of eligible publications increased over the past ten years. However, the methodological quality was critically low in as many as 90% of the studies retrieved and did not increase with time. The last authors H-index was the only quality predictor among the variables analysed. ConclusionThe methodological quality of the SR-MA of RCT is unacceptably low. Given the frequent application of physiotherapy in musculoskeletal disorders, there is an urgent need to improve secondary research by adopting more rigorous methods. RegistrationOpen Science Framework (https://osf.io/bc8zw/)

12
Does the type of publisher response to integrity concerns influence subsequent citations? A cohort study.

Studd, H.; Avenell, A.; Grey, A.; Bolland, M.

2026-02-27 health informatics 10.64898/2026.02.25.26346683 medRxiv
Top 0.1%
33.2%
Show abstract

BackgroundJournals may respond to integrity concerns by publishing an editorial response (editorial notice, expression of concern (EoC) or retraction). We investigated whether the type of editorial response affected citation rates. MethodsWe obtained citations for 172 randomised controlled trials (RCTs) with integrity concerns (41 had editorial notices, 38 EoCs and 23 retractions) and control RCTs from the same journal and year. Monthly citation rates up to 60 months before and after editorial responses were compared by editorial response type, and to citation decline in control RCTs. Results172 RCTs had 10,603 citations from 6,376 articles. 3,330 control trials were identified for 151/172 RCTs (15,948 citations, 87,811 articles). For both groups, citations increased steadily, peaking 45-65 months post-publication. There were no statistically significant differences in citation decline post-editorial response for trials receiving a retraction, EoC, or notice. Citations were lower in controls than index trials, so analyses were restricted to 1598 highly cited (>25) controls. The rate of decline for highly cited control trials was not statistically significantly different from the post-editorial response rate for index groups. ConclusionCitation rate decline after editorial responses did not differ by type of editorial response nor from the natural decline in control trials. HighlightsO_LIJournals may respond to integrity concerns by issuing an editorial notice. C_LIO_LIThe effect of expressions of concern or other editorial notices on citation patterns is unclear. C_LIO_LIEditorial notices did not accelerate citation decline compared with control trials. C_LIO_LIThe type of notice was not associated with differences in citation decline. C_LIO_LILate editorial notices appear ineffective in preventing continued citation. C_LI

13
Time-Varying Cardiovascular Risk of Febuxostat versus Allopurinol in Gout: A One-stage Meta-analysis

Zhang, S.

2025-12-30 rheumatology 10.64898/2025.12.29.25343180 medRxiv
Top 0.1%
31.3%
Show abstract

ObjectivesWhether febuxostat is associated with an increased cardiovascular risk compared to allopurinol in patients with gout remains controversial, with major randomized trials reporting conflicting results. This study aims to perform a comprehensive meta-analysis using reconstructed individual participant data (IPD) to evaluate the time-varying cardiovascular and mortality risk of febuxostat versus allopurinol in gout. MethodsWe conducted a one-stage individual participant data meta-analysis. PubMed, Embase, and Cochrane Library were searched up to November 2025 for randomized controlled trials and propensity score-matched observational studies reporting Kaplan-Meier curves for cardiovascular risk or all-cause mortality. Individual patient data were reconstructed from published curves. Time-varying hazard ratios (HRs) were estimated using a mixed-effects Cox model with treatment-by-time interaction across pre-specified intervals (0-12, 12-24, 24-48, >48 months). ResultsFour studies (n=19,090) were included. Analysis revealed significant time-varying effects. For cardiovascular risk, HRs were 0.89 (95% CI 0.79 to 0.99) at 0-12 months, 0.58 (0.50 to 0.68) at 12-24 months, 0.86 (0.74 to 0.99) at 24-48 months, and 1.09 (0.81 to 1.46) after >48 months. For all-cause mortality, HRs were 0.70 (0.62 to 0.79), 0.39 (0.33 to 0.45), 0.75 (0.65 to 0.86), and 1.02 (0.77 to 1.34) across the same intervals, respectively. ConclusionThe cardiovascular and mortality risk of febuxostat relative to allopurinol is time-dependent, showing a significant early reduction that attenuates over time. These findings advocate for a time-aware approach to clinical management and monitoring.

14
Science or Advocacy? The Global Rise of Policy Claims in Population Health Research (1990-2024)

Bann, D.; Wang, M.; Davies, N. M.; Wright, L.; O'Connor, M.; Courtin, E.

2025-11-15 epidemiology 10.1101/2025.11.13.25340175 medRxiv
Top 0.1%
31.1%
Show abstract

Should original research routinely contain prominent policy claims, such as recommendations for policymakers or broad calls to action? Growing emphasis on "research impact" might be welcome but also have unintended consequences that include risks of overextrapolation, the blurring of roles between scientists and advocates, and potential erosion of scientific credibility. To inform this debate, we examined 45,807 abstracts from ten leading Epidemiology and Public Health journals (1990-2024). Using a large language model with human validation, we classified policy claims and mapped their prevalence by time, country, journal, field of study, and study design. Policy claims markedly increased in frequency from 17.6% in 1990-1999 to 35.8% in 2020-2024, with wide variation across countries (>40% in Italy and Australia vs <19% in Norway and Japan, 2020-2024) and journals (>60% in some vs <6% in others). Keywords linked to higher claim rates differed by topic and time: some corresponded to topics with clear causal evidence ("tobacco"), others to topics with more complex causal evidence and notable researcher advocacy ("health inequalities" and "COVID-19"). Claims were most common in qualitative or cross-sectional studies, and less common in cohort, quasi-experimental, or experimental studies. We argue that these patterns reflect a research culture increasingly oriented toward claiming policy relevance--and incentives that encourage attaching claims to single studies. Our findings raise questions about how scientists and journals balance evidence, advocacy, and credibility. Ensuring that policy claims remain commensurate with evidence will be central to build trust as policy impact continues to be incentivised. We discuss alternative ways for researchers and publishers to engage meaningfully with policy beyond attaching claims to individual studies, and share our data and scripts to catalyse further work in this area.

15
Validation of an AI-Assisted Framework for Systematic Bias Assessment in Observational Studies

Etminan, M.; Rezaeianzadeh, R.; Douros, A.

2026-04-28 epidemiology 10.64898/2026.04.26.26351778 medRxiv
Top 0.1%
27.6%
Show abstract

BackgroundThe rapid expansion of medical literature has led to substantial variability and frequent contradictions in study findings, making it increasingly difficult to distinguish meaningful signals from noise. Much of this variability arises from differences in study methodology, where biases such as confounding, selection bias, and reverse causation can drive spurious associations. While artificial intelligence (AI)-assisted tools have been developed to support risk-of-bias assessment, most are designed for systematic reviews and are not tailored to identifying specific epidemiologic biases in observational studies. This highlights the need for structured, scalable approaches to evaluate study validity in real-world evidence. ObjectiveTo develop and validate an AI-assisted, expert-informed, rule-based framework (EpiVise) for systematically identifying and classifying key sources of bias in pharmacoepidemiologic studies, and to assess its agreement with expert evaluation. MethodsWe conducted a validation study using recently published pharmacoepidemiologic studies from high-impact journals (post-2025). Each study was independently assessed by the framework and two expert epidemiologists, across predefined bias domains, including measured confounding, confounding by indication, selection bias, immortal time bias, and disease latency. Agreement was evaluated using weighted kappa statistics. In the absence of a gold standard, expert judgment served as the reference benchmark. In a second phase, synthetic study scenarios with predefined embedded biases were constructed to assess the frameworks ability to detect known bias structures under controlled conditions. ResultsIn analyses of published studies (10 studies; 60 ratings), agreement between the framework and expert assessments was substantial ({kappa} = 0.75; 95% confidence interval [CI], 0.60-0.86), with 12 discordant ratings (20.0%), all limited to adjacent categories and occurring primarily in the confounding by indication and selection bias domains. In synthetic study scenarios (10 studies; 50 ratings), agreement was similarly substantial, with 42 of 50 ratings concordant (84%) and a weighted kappa of 0.77 (95% CI, 0.67-0.87); discordances included both adjacent-category and extreme disagreements and were concentrated in confounding by indication, selection bias, and prevalent user bias domains. ConclusionsThis AI-assisted, expert-informed framework, EpiVise provides a scalable and reproducible approach for evaluating epidemiologic study validity, substantial demonstrating agreement comparable to expert assessment. By systematically identifying key sources of bias, the framework has the potential to enhance the rigor and consistency of evidence evaluation, support peer review, and inform clinical, regulatory, and policy decision-making. Further validation across broader study designs and domains is warranted.

16
Estimating the replicability of highly cited clinical research (2004-2018)

da Costa, G. G.; Neves, K.; Amaral, O. B.

2022-05-31 epidemiology 10.1101/2022.05.31.22275810 medRxiv
Top 0.1%
27.6%
Show abstract

IntroductionPrevious studies about the replicability of clinical research based on the published literature have suggested that highly cited articles are often contradicted or found to have inflated effects. Nevertheless, there are no recent updates of such efforts, and this situation may have changed over time. MethodsWe searched the Web of Science database for articles studying medical interventions with more than 2000 citations, published between 2004 and 2018 in high-impact medical journals. We then searched for replications of these studies in PubMed using the PICO (Population, Intervention, Comparator and Outcome) framework. Replication success was evaluated by the presence of a statistically significant effect in the same direction and by overlap of the replications effect size confidence interval (CIs) with that of the original study. Evidence of effect size inflation and potential predictors of replicability were also analyzed. ResultsA total of 89 eligible studies, of which 24 had valid replications (17 meta-analyses and 7 primary studies) were found. Of these, 21 (88%) had effect sizes with overlapping CIs. Of 15 highly cited studies with a statistically significant difference in the primary outcome, 13 (87%) had a significant effect in the replication as well. When both criteria were considered together, the replicability rate in our sample was of 20 out of 24 (83%). There was no evidence of systematic inflation in these highly cited studies, with a mean effect size ratio of 1.03 (95% CI [0.88, 1.21]) between initial and subsequent effects. Due to the small number of contradicted results, our analysis had low statistical power to detect predictors of replicability. ConclusionAlthough most studies did not have eligible replications, the replicability rate of highly cited clinical studies in our sample was higher than in previous estimates, with little evidence of systematic effect size inflation.

17
Silence is golden, by my measures still see: why cheap-but-noisy outcome measures can be more cost effective than gold standards.

Woolf, B.; Pedder, H.; Rodriguez-Broadbent, H.; Edwards, P.

2022-05-19 epidemiology 10.1101/2022.05.17.22274839 medRxiv
Top 0.1%
26.2%
Show abstract

ObjectiveTo assess the cost-effectiveness of using cheap-but-noisy outcome measures, such as a short and simple questionnaire. BackgroundTo detect associations reliably, studies must avoid bias and random error. To reduce random error, we can increase the size of the study and increase the accuracy of the outcome measurement process. However, with fixed resources there is a trade-off between the number of participants a study can enrol and the amount of information that can be collected on each participant during data collection. MethodTo consider the effect on measurement error of using outcome scales with varying numbers of categories we define and calculate the Variance from Categorisation that would be expected from using a category midpoint; define the analytic conditions under-which such a measure is cost-effective; use meta-regression to estimate the impact of participant burden, defined as questionnaire length, on response rates; and develop an interactive web-app to allow researchers to explore the cost-effectiveness of using such a measure under plausible assumptions. ResultsCompared with no measurement, only having a few categories greatly reduced the Variance from Categorization. For example, scales with five categories reduce the variance by 96% for a uniform distribution. We additionally show that a simple measure will be more cost effective than a gold-standard measure if the relative increase in variance due to using it is less than the relative increase in cost from the gold standard, assuming it does not introduce bias in the measurement. We found an inverse power law relationship between participant burden and response rates such that a doubling the burden on participants reduces the response rate by around one third. Finally, we created an interactive web-app (https://benjiwoolf.shinyapps.io/cheapbutnoisymeasures/) to allow exploration of when using a cheap-but-noisy measure will be more cost-effective using realistic parameter. ConclusionCheap-but-noisy questionnaires containing just a few questions can be a cost effect way of maximising power. However, their use requires a judgment on the trade-off between the potential increase in risk information bias and the reduction in the potential of selection bias due to the expected higher response rates. Key Messages- A cheap-but-noisy outcome measure, like a short form questionnaire, is a more cost-effective method of maximising power than an error free gold standard when the percentage increase in noise from using the cheap-but-noisy measure is less than the relative difference in the cost of administering the two alternatives. - We have created an R-shiny app to facilitate the exploration of when this condition is met at https://benjiwoolf.shinyapps.io/cheapbutnoisymeasures/ - Cheap-but-noisy outcome measures are more likely to introduce information bias than a gold standard, but may reduce selection bias because they reduce loss-to-follow-up. Researchers therefore need to form a judgement about the relative increase or decrease in bias before using a cheap-but-noisy measure. - We would encourage the development and validation of short form questionnaires to enable the use of high quality cheap-but-noisy outcome measures in randomised controlled trials.

18
Prevalence of bias attributable to composite outcome in clinical trials: a systematic review

Silva, J. M. N. d.; Conceicao, J. F. S.; Ramirez, P. C.; Diaz-Leon, C. L.; Diaz-Quijano, F. A.

2024-07-18 epidemiology 10.1101/2024.07.18.24310633 medRxiv
Top 0.1%
23.3%
Show abstract

ObjectiveTo investigate the prevalence of bias attributable to composite outcome (BACO) in clinical trials. Study design and settingWe searched PubMed for randomized clinical trials where the primary outcome was a binary composite that included all-cause mortality among its components from January 1, 2019, to December 31, 2020. For each trial, the BACO index was calculated to assess the correspondence between effects on the composite outcome and that on mortality. This systematic review was registered in PROSPERO (CRD42021229554). ResultsAfter screening 1,076 citations and 171 full-text articles, 91 studies were included from 13 different medical areas. The prevalence of significant or suggestive BACO among the 91 included articles was 25.2% (n=23), including 12 with p<0.005 and 11 with p between 0.005 and <0.05. We observed that in 17 (73.9%) of these 23 studies, the BACO index value was between zero and <1, indicating an underestimation of the effect. The other six studies showed negative values (26.1%), indicating an inversion of the association with mortality. None of the studies showed significant overestimation of the association attributable to the composite outcome. ConclusionThese findings highlight the need to predefine guidelines for interpreting effects on composite endpoints based on objective criteria such as the BACO index. What is new?O_ST_ABSKey FindingsC_ST_ABSO_LIThe study found that 25.2% of the included clinical trials exhibited significant or suggestive bias attributable to composite outcomes (BACO). C_LIO_LIIn 73.9% of these cases, the BACO index was less than 1, indicating an underestimation of the effect. 26.1% of the studies showed an inversion of the association with mortality. C_LIO_LINo significant overestimation of the association due to composite outcomes was observed. C_LI What This Adds to What Was Known?O_LIThis study contributes to the existing knowledge by quantifying the prevalence of bias attributable to composite outcomes in clinical trials. C_LIO_LIIt highlights that a significant proportion of trials may underestimate the effect or even show an inversion of the association with mortality when composite outcomes are used. C_LIO_LIThis finding emphasizes the need for careful consideration and objective criteria, like the BACO index, in the design and interpretation of clinical trials involving composite outcomes. C_LI What Is the Implication and What Should Change Now?O_LIResearchers and clinicians should be cautious about relying solely on composite outcomes without assessing the potential biases they introduce. C_LIO_LIThe study suggests a need for predefined guidelines and objective criteria, such as the BACO index, for interpreting the effects of composite outcomes. C_LI

19
Registration and reporting characteristics of trials investigating exercise therapy following total knee arthroplasty: A systematic review.

Groenfeldt, B. M.; Husted, R. S.; Broedsgaard, R. H.; Holst, L.; Kallemose, T.; Chapple, C.; Juhl, C. B.; Bandholm, T.

2025-10-08 rehabilitation medicine and physical therapy 10.1101/2025.10.07.25337493 medRxiv
Top 0.1%
23.0%
Show abstract

ObjectivesProspectively registering the primary trial outcome is important to reduce selective outcome reporting and increase trustworthiness of findings used to guide clinical practice. The objective of this systematic review was to explore and compare the reporting characteristics of prospectively and non-prospectively registered trials investigating exercise therapy following total knee arthroplasty. DesignTrials comparing effects of exercise therapy after total knee arthroplasty due to osteoarthritis were sought in four databases from 2000 (clinicaltrials.gov launch) to 12th of August 2024. Randomised controlled trials comparing different exercise therapy interventions were included. Primary outcomes were extracted using a pre-specified hierarchy to reflect each trials most consistently reported outcome. Risk-of-bias was assessed using the Cochrane Collaborations Risk-of-Bias tool version 2. ResultsNinety-four trials (n = 9,396) were included: 13 prospectively registered, 33 retrospectively registered, and 48 unregistered. A single primary outcome was defined in 43.6% of trials. Four trials reported a primary outcome consistent with a prospective registration. Prospectively registered trials reported smaller effect estimates (SMD 0.06; 95% CI -0.03 to 0.16) than retrospectively registered (SMD 0.67; 95% CI 0.22 to 1.11) and unregistered trials (SMD 0.59; 95% CI 0.32 to 0.86), and more often defined a single primary outcome, reported sample size calculations, and followed intention-to-treat principles. ConclusionProspective registration and clear primary outcome definition were uncommon in trials comparing exercise therapy after total knee arthroplasty. Trials without these elements had larger effect size estimates at higher risk of bias, suggesting that registration status is an important indicator of methodological quality. Registrationhttps://doi.org/10.17605/OSF.IO/6WZRE Protocol and SAPhttps://osf.io/kwqcb/files/osfstorage

20
Research waste from poor reporting of core methods and results and redundancy in studies of reporting guideline adherence: a meta-research review

Dal Santo, T.; Rice, D. B.; Amiri, L. S.; Tasleem, A.; Li, K.; Boruff, J. T.; Geoffroy, M.-C. B.; Benedetti, A.; Thombs, B.

2022-12-20 epidemiology 10.1101/2022.12.19.22283669 medRxiv
Top 0.1%
22.9%
Show abstract

ObjectivesWe investigated meta-research studies that evaluated adherence to prominent reporting guidelines (CONSORT, PRISMA, STARD, STROBE) in health research studies to determine the proportion that (1) provided an explanation for how complex guideline items were rated for adherence and (2) provided results from individual studies reviewed in addition to aggregate results. We also examined the conclusions of each meta-research study to assess redundancy of findings across studies. DesignCross-sectional meta-research review. Data sourcesMEDLINE (Ovid) searched on July 5, 2022. Eligibility criteria for selecting studiesStudies in any language were eligible if they used any version of the CONSORT, PRISMA, STARD, or STROBE reporting guidelines or their extensions to evaluate reporting in at least 10 human health research studies. We excluded studies that modified a reporting guideline or its items or evaluated fewer than half of reporting guideline items. Main outcomes were (1) the proportion of meta-research studies that provided a coding explanation that could be used to replicate the study or verify its results and (2) the proportion that provided individual-level study results in the main text, supplemental materials, or via an internet link. ResultsOf 148 included meta-research studies, 14 (10%, 95% confidence interval [CI] 6% to 15%) provided a fully replicable coding explanation, and 49 (33%, 95% CI 26% to 41%) completely reported individual study results. Of 90 studies that classified reporting as adequate or inadequate in the study abstract, 6 (7%, 95% CI 3% to 14%) concluded that reporting was adequate but none of those 6 studies provided information on how items were coded or provided item-level results for included studies. ConclusionsMuch of published meta-research on reporting in health research is likely wasteful. Few studies report enough information for verification or replication, and almost all find that reporting in health research studies is suboptimal. These findings highlight the importance of shifting the focus from assessing reporting adequacy to developing, testing, and implementing strategies to improve reporting. FundingThere was no specific funding for this study. ProtocolPosted on the Open Science Framework June 29, 2022 (https://osf.io/gtm4z/).